Skip to content

Replace hash groupby internals with HashCSR - #24050

Draft
PointKernel wants to merge 9 commits into
NVIDIA:mainfrom
PointKernel:hashcsr-groupby
Draft

Replace hash groupby internals with HashCSR#24050
PointKernel wants to merge 9 commits into
NVIDIA:mainfrom
PointKernel:hashcsr-groupby

Conversation

@PointKernel

@PointKernel PointKernel commented Sep 9, 2026

Copy link
Copy Markdown
Member

Description

This PR replaces the hash groupby internals with a HashCSR build followed by CUB reductions, in the spirit of #23640 for hash join. Rows are grouped by a single probe pass that gives every row a slot and a rank, the occupied slots are compacted into groups, and a fill pass produces the grouped row order. For inputs of two million rows or more with keys up to 32 bytes, the table is sized from a sampled estimate of the number of distinct keys rather than the row count, with a bounded probe that falls back to a slot per row when the estimate is short; this keeps the table cache-resident and roughly halves peak memory on low- and mid-cardinality inputs.

Aggregations are DeviceSegmentedReduce calls over the grouped order. Groups averaging under 128 rows are packed several per block through the dispatch's segment-size hint (one thread or one 8-lane sub-warp per group), larger groups get a block each, and groups longer than one chunk are reduced in two levels. Nullable aggregations carry the group validity in the accumulator, so a result and its null mask come out of one pass, and the sums behind MEAN, M2, VARIANCE and STD are fused into one reduction.

This removes the shared-memory aggregation kernel, the block-local mapping kernel, and the dense and sparse global-memory atomic paths (about 600 lines net). All dispatch is on the host, and because reductions no longer need atomics, decimal128 MIN/MAX and fixed-point SUM_OVERFLOW now take the hash path instead of the sort-based one.

Checklist

  • I am familiar with the Contributing Guidelines.
  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

@PointKernel
PointKernel requested review from a team as code owners September 9, 2026 02:34
@PointKernel
PointKernel requested a review from ttnghia September 9, 2026 02:34
@PointKernel PointKernel added libcudf Affects libcudf (C++/CUDA) code. Performance Performance related issue improvement Improvement / enhancement to an existing function non-breaking Non-breaking change labels Sep 9, 2026
@PointKernel
PointKernel requested a review from lamarrr September 9, 2026 02:34
@github-actions github-actions Bot added the CMake CMake build issue label Sep 9, 2026
@PointKernel
PointKernel marked this pull request as draft September 9, 2026 02:35
@copy-pr-bot

copy-pr-bot Bot commented Sep 9, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@PointKernel PointKernel added the 2 - In Progress Currently a work in progress label Sep 9, 2026
@coderabbitai

coderabbitai Bot commented Sep 9, 2026

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: dc27a7c7-1ed1-40e6-83b0-4895912f9e81

📥 Commits

Reviewing files that changed from the base of the PR and between e44e3cf and 67c719a.

📒 Files selected for processing (24)
  • cpp/CMakeLists.txt
  • cpp/src/groupby/hash/compute_global_memory_aggs.cu
  • cpp/src/groupby/hash/compute_global_memory_aggs.cuh
  • cpp/src/groupby/hash/compute_global_memory_aggs.hpp
  • cpp/src/groupby/hash/compute_global_memory_aggs_null.cu
  • cpp/src/groupby/hash/compute_groupby.cu
  • cpp/src/groupby/hash/compute_mapping_indices.cu
  • cpp/src/groupby/hash/compute_mapping_indices.cuh
  • cpp/src/groupby/hash/compute_mapping_indices.hpp
  • cpp/src/groupby/hash/compute_mapping_indices_null.cu
  • cpp/src/groupby/hash/compute_shared_memory_aggs.cu
  • cpp/src/groupby/hash/compute_shared_memory_aggs.hpp
  • cpp/src/groupby/hash/compute_single_pass_aggs.cu
  • cpp/src/groupby/hash/compute_single_pass_aggs.cuh
  • cpp/src/groupby/hash/compute_single_pass_aggs.hpp
  • cpp/src/groupby/hash/compute_single_pass_aggs_null.cu
  • cpp/src/groupby/hash/global_memory_aggregator.cuh
  • cpp/src/groupby/hash/groupby.cu
  • cpp/src/groupby/hash/hash_csr_kernels.cuh
  • cpp/src/groupby/hash/helpers.cuh
  • cpp/src/groupby/hash/output_utils.cu
  • cpp/src/groupby/hash/output_utils.hpp
  • cpp/src/groupby/hash/shared_memory_aggregator.cuh
  • cpp/src/groupby/hash/single_pass_functors.cuh
💤 Files with no reviewable changes (17)
  • cpp/src/groupby/hash/compute_mapping_indices_null.cu
  • cpp/src/groupby/hash/compute_global_memory_aggs.cuh
  • cpp/src/groupby/hash/compute_global_memory_aggs.cu
  • cpp/src/groupby/hash/compute_global_memory_aggs.hpp
  • cpp/src/groupby/hash/shared_memory_aggregator.cuh
  • cpp/src/groupby/hash/global_memory_aggregator.cuh
  • cpp/src/groupby/hash/compute_global_memory_aggs_null.cu
  • cpp/src/groupby/hash/compute_mapping_indices.hpp
  • cpp/src/groupby/hash/compute_mapping_indices.cuh
  • cpp/src/groupby/hash/compute_shared_memory_aggs.cu
  • cpp/CMakeLists.txt
  • cpp/src/groupby/hash/compute_shared_memory_aggs.hpp
  • cpp/src/groupby/hash/compute_single_pass_aggs.cuh
  • cpp/src/groupby/hash/compute_mapping_indices.cu
  • cpp/src/groupby/hash/output_utils.hpp
  • cpp/src/groupby/hash/compute_single_pass_aggs_null.cu
  • cpp/src/groupby/hash/output_utils.cu

Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Summary

Summary by CodeRabbit

  • New Features
    • Updated hash-based groupby processing with improved grouped-row handling and key organization.
    • Expanded single-pass aggregation support, including counts, sums, products, minimums, maximums, argument-based aggregates, and sum-of-squares.
    • Improved handling of nullable data and aggregation result validity.
    • Added support for combining compatible consecutive aggregations during groupby processing.
  • Bug Fixes
    • Added validation for groupby capacity, input shapes, and supported aggregation types.

Walkthrough

Hash groupby now uses HashCSR to create grouped rows and a new single-pass aggregation pipeline. Obsolete cuco-based grouping, shared-memory, global-memory, mapping-index, and sparse-output implementations are removed.

Changes

Hash groupby migration

Layer / File(s) Summary
HashCSR grouping pipeline
cpp/src/groupby/hash/hash_csr_kernels.cuh, cpp/src/groupby/hash/compute_groupby.cu
HashCSR builds grouped keys, group offsets, representative rows, and grouped input rows for aggregation.
Grouped-row aggregation pipeline
cpp/src/groupby/hash/compute_single_pass_aggs.hpp, cpp/src/groupby/hash/compute_single_pass_aggs.cu
Single-pass aggregation now uses grouped rows, adaptive chunking, typed reductions, null handling, fused additive operations, and a column-based API.
Aggregation capability wiring
cpp/src/groupby/hash/groupby.cu
Hash groupby eligibility now delegates to is_single_pass_agg_supported.
Legacy hash paths removal
cpp/CMakeLists.txt, cpp/src/groupby/hash/*
Obsolete cuco-based grouping, shared-memory and global-memory aggregation, mapping-index, sparse-output, and key-transformation code is deleted.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~60 minutes

Merge Risk: ⚪ Minimal · up to 67c71

No concrete current-head issue remains from the finalized findings; the PR is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 5.56% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 36 functions across 4 files. (3 skipped: 3… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely summarizes the primary change: replacing the hash groupby internals with HashCSR.
Description check ✅ Passed The description directly explains the HashCSR groupby implementation, segmented reductions, removed aggregation paths, and resulting behavior changes.
Full details: Docstring Coverage

Explanation

Docstring coverage is 5.56% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 36 functions across 4 files. (3 skipped: 3 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@PointKernel

Copy link
Copy Markdown
Member Author

Groupby benchmarks vs main, RTX PRO 6000 (sm_120), CUDA 12.9, idle GPU. nvbench_compare.py output, status relative to measured noise. Extra axes added for 1 to 8 groups, 200K to 2M groups, and 32 to 200 rows per group.

benchmark faster slower same
groupby_max_cardinality (20M rows, MAX) 56 0 0
groupby_max_cardinality, 1 to 8 groups 0 6 2
groupby_max_cardinality, 200K to 2M groups 4 2 0
complex_int_keys 20 0 0
complex_mixed_keys 36 4 0
groupby_max 35 1 0
groupby_struct_keys 18 0 0
groupby_m2_var_std 42 6 0
groupby_m2_var_std, 32 to 200 rows per group 13 7 0
sum 6 2 0
no_requests 3 0 1
total 233 28 3

Slower cases: 4 to 8 groups with 8 aggregations, variance on ~20-row groups with 50% nulls, a single sum at 100M rows, and string keys at 84K groups without nulls. Peak memory is unchanged for distinct keys and 15 to 40% higher when aggregating. decimal128 MIN/MAX moves from the sort path to hash (12x to 91x).

groupby_max_cardinality (20M rows, 3 int32 key columns, MAX): 56 faster, 0 slower, 0 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 20 normal 2.791 ms 5.37% 1.044 ms 0.71% -1746.657 us -62.59% 🟢 FAST
I32 20000000 2 20 normal 3.386 ms 76.79% 1.204 ms 0.55% -2181.493 us -64.43% 🟢 FAST
I32 20000000 4 20 normal 2.563 ms 1.18% 1.532 ms 0.47% -1030.455 us -40.21% 🟢 FAST
I32 20000000 8 20 normal 2.956 ms 1.20% 2.182 ms 0.44% -773.879 us -26.18% 🟢 FAST
I32 20000000 1 50 normal 2.316 ms 1.02% 1.056 ms 0.77% -1259.631 us -54.40% 🟢 FAST
I32 20000000 2 50 normal 2.514 ms 2.42% 1.225 ms 0.76% -1288.773 us -51.26% 🟢 FAST
I32 20000000 4 50 normal 2.788 ms 2.14% 1.567 ms 0.50% -1220.241 us -43.77% 🟢 FAST
I32 20000000 8 50 normal 3.216 ms 1.96% 2.230 ms 0.30% -985.850 us -30.65% 🟢 FAST
I32 20000000 1 100 normal 2.522 ms 2.28% 1.051 ms 0.43% -1470.992 us -58.33% 🟢 FAST
I32 20000000 2 100 normal 2.692 ms 2.12% 1.223 ms 0.40% -1469.086 us -54.58% 🟢 FAST
I32 20000000 4 100 normal 2.949 ms 2.01% 1.571 ms 0.57% -1378.581 us -46.74% 🟢 FAST
I32 20000000 8 100 normal 3.413 ms 1.83% 2.248 ms 0.45% -1164.403 us -34.12% 🟢 FAST
I32 20000000 1 1000 normal 4.116 ms 1.90% 1.054 ms 0.49% -3062.604 us -74.40% 🟢 FAST
I32 20000000 2 1000 normal 4.601 ms 3.19% 1.228 ms 0.43% -3372.791 us -73.31% 🟢 FAST
I32 20000000 4 1000 normal 8.817 ms 0.77% 1.558 ms 0.32% -7259.532 us -82.33% 🟢 FAST
I32 20000000 8 1000 normal 13.956 ms 0.59% 2.223 ms 0.27% -11732.861 us -84.07% 🟢 FAST
I32 20000000 1 10000 normal 4.036 ms 1.57% 1.075 ms 0.45% -2961.148 us -73.37% 🟢 FAST
I32 20000000 2 10000 normal 4.278 ms 1.57% 1.243 ms 0.44% -3035.113 us -70.95% 🟢 FAST
I32 20000000 4 10000 normal 4.949 ms 1.20% 1.568 ms 0.32% -3381.634 us -68.32% 🟢 FAST
I32 20000000 8 10000 normal 7.887 ms 37.03% 2.228 ms 0.20% -5658.653 us -71.75% 🟢 FAST
I32 20000000 1 100000 normal 4.842 ms 6.23% 1.244 ms 0.30% -3598.867 us -74.32% 🟢 FAST
I32 20000000 2 100000 normal 5.142 ms 5.85% 1.450 ms 0.25% -3691.873 us -71.80% 🟢 FAST
I32 20000000 4 100000 normal 5.918 ms 46.94% 1.846 ms 0.49% -4071.837 us -68.80% 🟢 FAST
I32 20000000 8 100000 normal 6.172 ms 5.02% 2.644 ms 0.13% -3527.984 us -57.16% 🟢 FAST
I32 20000000 1 1000000 normal 12.197 ms 0.10% 2.870 ms 0.19% -9326.107 us -76.46% 🟢 FAST
I32 20000000 2 1000000 normal 17.117 ms 6.04% 3.183 ms 0.17% -13933.926 us -81.40% 🟢 FAST
I32 20000000 4 1000000 normal 15.483 ms 6.95% 3.827 ms 0.50% -11656.545 us -75.29% 🟢 FAST
I32 20000000 8 1000000 normal 9.383 ms 47.69% 5.110 ms 0.14% -4273.653 us -45.55% 🟢 FAST
decimal128 20000000 1 20 normal 96.664 ms 0.07% 1.260 ms 0.76% -95404.035 us -98.70% 🟢 FAST
decimal128 20000000 2 20 normal 97.978 ms 0.21% 1.622 ms 0.35% -96356.132 us -98.34% 🟢 FAST
decimal128 20000000 4 20 normal 98.065 ms 0.34% 2.363 ms 0.41% -95701.571 us -97.59% 🟢 FAST
decimal128 20000000 8 20 normal 104.450 ms 0.21% 3.891 ms 0.18% -100558.573 us -96.27% 🟢 FAST
decimal128 20000000 1 50 normal 100.640 ms 0.05% 1.255 ms 0.50% -99384.552 us -98.75% 🟢 FAST
decimal128 20000000 2 50 normal 101.800 ms 0.25% 1.647 ms 0.25% -100153.816 us -98.38% 🟢 FAST
decimal128 20000000 4 50 normal 103.609 ms 0.06% 2.417 ms 0.20% -101191.207 us -97.67% 🟢 FAST
decimal128 20000000 8 50 normal 108.919 ms 0.03% 3.955 ms 0.14% -104964.937 us -96.37% 🟢 FAST
decimal128 20000000 1 100 normal 104.253 ms 0.27% 1.261 ms 0.33% -102992.462 us -98.79% 🟢 FAST
decimal128 20000000 2 100 normal 150.131 ms 9.14% 1.658 ms 0.35% -148473.577 us -98.90% 🟢 FAST
decimal128 20000000 4 100 normal 115.670 ms 19.20% 2.441 ms 0.23% -113228.626 us -97.89% 🟢 FAST
decimal128 20000000 8 100 normal 106.146 ms 0.17% 3.999 ms 0.16% -102147.753 us -96.23% 🟢 FAST
decimal128 20000000 1 1000 normal 106.161 ms 0.14% 1.274 ms 0.31% -104887.785 us -98.80% 🟢 FAST
decimal128 20000000 2 1000 normal 107.257 ms 0.12% 1.667 ms 0.26% -105589.245 us -98.45% 🟢 FAST
decimal128 20000000 4 1000 normal 109.292 ms 0.15% 2.447 ms 0.23% -106845.040 us -97.76% 🟢 FAST
decimal128 20000000 8 1000 normal 113.612 ms 0.13% 3.998 ms 0.16% -109614.494 us -96.48% 🟢 FAST
decimal128 20000000 1 10000 normal 108.298 ms 0.12% 1.298 ms 0.42% -106999.987 us -98.80% 🟢 FAST
decimal128 20000000 2 10000 normal 109.306 ms 0.11% 1.676 ms 0.28% -107629.348 us -98.47% 🟢 FAST
decimal128 20000000 4 10000 normal 111.313 ms 0.10% 2.432 ms 0.21% -108881.649 us -97.82% 🟢 FAST
decimal128 20000000 8 10000 normal 115.562 ms 0.10% 3.948 ms 0.50% -111614.763 us -96.58% 🟢 FAST
decimal128 20000000 1 100000 normal 100.551 ms 0.06% 1.420 ms 0.30% -99130.772 us -98.59% 🟢 FAST
decimal128 20000000 2 100000 normal 101.673 ms 0.13% 1.768 ms 0.29% -99904.943 us -98.26% 🟢 FAST
decimal128 20000000 4 100000 normal 103.683 ms 0.15% 2.462 ms 0.34% -101221.837 us -97.63% 🟢 FAST
decimal128 20000000 8 100000 normal 107.909 ms 0.12% 3.861 ms 0.50% -104048.254 us -96.42% 🟢 FAST
decimal128 20000000 1 1000000 normal 90.185 ms 0.23% 3.216 ms 0.18% -86969.063 us -96.43% 🟢 FAST
decimal128 20000000 2 1000000 normal 91.259 ms 0.15% 3.881 ms 0.14% -87377.666 us -95.75% 🟢 FAST
decimal128 20000000 4 1000000 normal 93.385 ms 0.13% 5.214 ms 0.11% -88170.706 us -94.42% 🟢 FAST
decimal128 20000000 8 1000000 normal 97.414 ms 0.19% 7.894 ms 0.07% -89520.165 us -91.90% 🟢 FAST
groupby_max_cardinality, extra axis: 1 to 8 groups: 0 faster, 6 slower, 2 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 1 normal 1.029 ms 0.95% 1.030 ms 0.61% 1.293 us 0.13% 🔵 SAME
I32 20000000 8 1 normal 1.661 ms 0.96% 1.741 ms 0.35% 80.285 us 4.83% 🔴 SLOW
I32 20000000 1 2 normal 1.024 ms 0.95% 1.027 ms 0.43% 2.741 us 0.27% 🔵 SAME
I32 20000000 8 2 normal 1.690 ms 1.23% 1.830 ms 0.25% 139.725 us 8.27% 🔴 SLOW
I32 20000000 1 4 normal 1.024 ms 0.95% 1.214 ms 0.95% 189.459 us 18.50% 🔴 SLOW
I32 20000000 8 4 normal 1.673 ms 1.62% 2.157 ms 1.00% 484.319 us 28.96% 🔴 SLOW
I32 20000000 1 8 normal 1.023 ms 0.88% 1.128 ms 0.44% 104.572 us 10.22% 🔴 SLOW
I32 20000000 8 8 normal 1.718 ms 1.82% 2.116 ms 0.24% 398.483 us 23.20% 🔴 SLOW
groupby_max_cardinality, extra axis: 200K to 2M groups: 4 faster, 2 slower, 0 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 200000 normal 2.769 ms 0.32% 1.517 ms 0.25% -1252.467 us -45.22% 🟢 FAST
I32 20000000 8 200000 normal 4.219 ms 0.13% 3.668 ms 0.10% -550.670 us -13.05% 🟢 FAST
I32 20000000 1 500000 normal 3.397 ms 0.50% 2.214 ms 0.17% -1183.015 us -34.83% 🟢 FAST
I32 20000000 8 500000 normal 4.415 ms 0.15% 4.429 ms 0.14% 14.168 us 0.32% 🔴 SLOW
I32 20000000 1 2000000 normal 4.314 ms 0.24% 3.421 ms 0.14% -893.283 us -20.71% 🟢 FAST
I32 20000000 8 2000000 normal 5.271 ms 0.15% 5.689 ms 0.24% 417.840 us 7.93% 🔴 SLOW
complex_int_keys: 20 faster, 0 slower, 0 same
num_cols num_rows value_key_ratio null_probability Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
1 2^24 20 0 9.076 ms 15.44% 2.315 ms 0.23% -6760.679 us -74.49% 🟢 FAST
2 2^24 20 0 9.510 ms 14.72% 2.440 ms 0.27% -7070.671 us -74.35% 🟢 FAST
4 2^24 20 0 10.254 ms 13.44% 2.798 ms 0.19% -7456.233 us -72.71% 🟢 FAST
8 2^24 20 0 12.719 ms 9.46% 3.704 ms 0.17% -9014.795 us -70.88% 🟢 FAST
16 2^24 20 0 19.306 ms 5.88% 5.632 ms 0.50% -13673.658 us -70.83% 🟢 FAST
1 2^24 200 0 6.834 ms 21.59% 1.008 ms 0.44% -5825.521 us -85.25% 🟢 FAST
2 2^24 200 0 7.028 ms 20.59% 1.104 ms 0.50% -5924.331 us -84.29% 🟢 FAST
4 2^24 200 0 8.547 ms 16.40% 1.282 ms 0.35% -7265.254 us -85.00% 🟢 FAST
8 2^24 200 0 10.995 ms 13.41% 1.644 ms 0.44% -9350.441 us -85.05% 🟢 FAST
16 2^24 200 0 16.202 ms 7.23% 2.473 ms 0.23% -13729.231 us -84.74% 🟢 FAST
1 2^24 20 0.5 8.177 ms 44.51% 1.321 ms 0.36% -6855.913 us -83.85% 🟢 FAST
2 2^24 20 0.5 6.776 ms 33.91% 1.500 ms 0.64% -5276.160 us -77.87% 🟢 FAST
4 2^24 20 0.5 7.657 ms 3.37% 1.851 ms 0.37% -5805.835 us -75.82% 🟢 FAST
8 2^24 20 0.5 10.229 ms 28.40% 2.641 ms 0.56% -7588.064 us -74.18% 🟢 FAST
16 2^24 20 0.5 18.248 ms 0.23% 4.282 ms 0.24% -13966.260 us -76.54% 🟢 FAST
1 2^24 200 0.5 6.255 ms 41.95% 931.814 us 0.58% -5322.754 us -85.10% 🟢 FAST
2 2^24 200 0.5 5.382 ms 0.14% 1.124 ms 0.57% -4258.016 us -79.12% 🟢 FAST
4 2^24 200 0.5 4.811 ms 49.69% 1.396 ms 0.50% -3414.811 us -70.98% 🟢 FAST
8 2^24 200 0.5 4.709 ms 0.24% 1.955 ms 0.34% -2754.041 us -58.49% 🟢 FAST
16 2^24 200 0.5 8.314 ms 0.17% 3.148 ms 0.24% -5166.338 us -62.14% 🟢 FAST
complex_mixed_keys: 36 faster, 4 slower, 0 same
num_cols num_rows value_key_ratio null_probability Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
1 2^18 20 0 1.488 ms 4.52% 111.094 us 4.43% -1377.405 us -92.54% 🟢 FAST
2 2^18 20 0 3.066 ms 98.42% 208.925 us 3.16% -2857.116 us -93.19% 🟢 FAST
3 2^18 20 0 2.031 ms 2.03% 221.668 us 2.95% -1809.598 us -89.09% 🟢 FAST
4 2^18 20 0 6.931 ms 0.08% 2.698 ms 0.36% -4233.616 us -61.08% 🟢 FAST
5 2^18 20 0 7.525 ms 0.24% 2.796 ms 0.43% -4728.286 us -62.84% 🟢 FAST
1 2^24 20 0 5.840 ms 54.06% 2.317 ms 0.16% -3522.190 us -60.32% 🟢 FAST
2 2^24 20 0 8.812 ms 0.28% 7.020 ms 0.12% -1792.735 us -20.34% 🟢 FAST
3 2^24 20 0 10.362 ms 2.00% 9.228 ms 0.14% -1134.130 us -10.94% 🟢 FAST
4 2^24 20 0 47.576 ms 0.24% 43.615 ms 0.16% -3961.467 us -8.33% 🟢 FAST
5 2^24 20 0 54.229 ms 0.21% 48.281 ms 0.19% -5947.396 us -10.97% 🟢 FAST
1 2^18 200 0 1.243 ms 1.26% 107.001 us 4.80% -1136.415 us -91.39% 🟢 FAST
2 2^18 200 0 1.029 ms 79.08% 208.716 us 2.74% -819.853 us -79.71% 🟢 FAST
3 2^18 200 0 302.284 us 2.57% 222.015 us 2.67% -80.269 us -26.55% 🟢 FAST
4 2^18 200 0 3.134 ms 0.25% 2.691 ms 0.36% -443.601 us -14.15% 🟢 FAST
5 2^18 200 0 3.295 ms 0.34% 2.807 ms 0.38% -487.904 us -14.81% 🟢 FAST
1 2^24 200 0 1.909 ms 0.43% 1.008 ms 0.36% -901.245 us -47.21% 🟢 FAST
2 2^24 200 0 6.711 ms 0.50% 7.014 ms 0.14% 303.230 us 4.52% 🔴 SLOW
3 2^24 200 0 8.378 ms 0.24% 9.235 ms 0.13% 857.363 us 10.23% 🔴 SLOW
4 2^24 200 0 42.088 ms 0.13% 43.661 ms 0.09% 1.573 ms 3.74% 🔴 SLOW
5 2^24 200 0 48.008 ms 0.15% 48.171 ms 0.25% 162.300 us 0.34% 🔴 SLOW
1 2^18 20 0.5 217.288 us 2.92% 184.382 us 2.87% -32.906 us -15.14% 🟢 FAST
2 2^18 20 0.5 348.870 us 3.46% 313.863 us 2.40% -35.007 us -10.03% 🟢 FAST
3 2^18 20 0.5 348.334 us 2.83% 321.784 us 4.22% -26.549 us -7.62% 🟢 FAST
4 2^18 20 0.5 630.493 us 3.01% 554.922 us 4.92% -75.570 us -11.99% 🟢 FAST
5 2^18 20 0.5 1.178 ms 2.93% 1.085 ms 1.49% -92.371 us -7.84% 🟢 FAST
1 2^24 20 0.5 2.434 ms 0.69% 1.338 ms 0.43% -1096.389 us -45.04% 🟢 FAST
2 2^24 20 0.5 3.303 ms 0.43% 2.326 ms 0.33% -977.619 us -29.59% 🟢 FAST
3 2^24 20 0.5 2.687 ms 0.46% 1.918 ms 0.42% -769.494 us -28.63% 🟢 FAST
4 2^24 20 0.5 3.081 ms 0.36% 2.800 ms 0.45% -280.824 us -9.11% 🟢 FAST
5 2^24 20 0.5 3.437 ms 0.45% 3.135 ms 0.63% -302.283 us -8.79% 🟢 FAST
1 2^18 200 0.5 216.153 us 2.52% 173.612 us 3.61% -42.540 us -19.68% 🟢 FAST
2 2^18 200 0.5 347.703 us 2.26% 314.724 us 2.52% -32.979 us -9.48% 🟢 FAST
3 2^18 200 0.5 350.044 us 2.11% 321.588 us 2.12% -28.456 us -8.13% 🟢 FAST
4 2^18 200 0.5 634.212 us 2.17% 551.958 us 1.72% -82.254 us -12.97% 🟢 FAST
5 2^18 200 0.5 1.191 ms 1.56% 1.084 ms 2.02% -107.809 us -9.05% 🟢 FAST
1 2^24 200 0.5 1.881 ms 0.78% 945.177 us 0.56% -936.180 us -49.76% 🟢 FAST
2 2^24 200 0.5 3.318 ms 0.53% 2.323 ms 0.50% -994.876 us -29.98% 🟢 FAST
3 2^24 200 0.5 2.671 ms 0.73% 1.912 ms 0.35% -758.837 us -28.41% 🟢 FAST
4 2^24 200 0.5 3.072 ms 0.41% 2.800 ms 0.50% -272.372 us -8.87% 🟢 FAST
5 2^24 200 0.5 3.438 ms 0.60% 3.115 ms 0.39% -322.181 us -9.37% 🟢 FAST
groupby_max: 35 faster, 1 slower, 0 same
T cardinality num_rows null_probability num_aggregations Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 0 2^24 0 1 4.942 ms 1.44% 3.245 ms 0.18% -1697.169 us -34.34% 🟢 FAST
I32 0 2^24 0.1 1 5.679 ms 1.44% 3.453 ms 0.44% -2225.999 us -39.20% 🟢 FAST
I32 0 2^24 0.9 1 5.316 ms 1.57% 3.261 ms 0.13% -2054.982 us -38.66% 🟢 FAST
I32 0 2^24 0 2 5.640 ms 1.41% 3.511 ms 0.14% -2128.244 us -37.74% 🟢 FAST
I32 0 2^24 0.1 2 6.416 ms 1.28% 3.993 ms 0.18% -2423.081 us -37.76% 🟢 FAST
I32 0 2^24 0.9 2 5.681 ms 1.41% 3.689 ms 0.25% -1991.790 us -35.06% 🟢 FAST
I32 0 2^24 0 4 5.456 ms 1.37% 4.126 ms 0.14% -1330.495 us -24.38% 🟢 FAST
I32 0 2^24 0.1 4 6.130 ms 1.29% 4.839 ms 0.14% -1290.230 us -21.05% 🟢 FAST
I32 0 2^24 0.9 4 5.406 ms 1.52% 4.246 ms 0.12% -1160.477 us -21.46% 🟢 FAST
I32 0 2^24 0 8 6.265 ms 1.17% 5.138 ms 0.23% -1126.978 us -17.99% 🟢 FAST
I32 0 2^24 0.1 8 9.491 ms 42.20% 6.816 ms 0.50% -2674.488 us -28.18% 🟢 FAST
I32 0 2^24 0.9 8 23.600 ms 9.31% 5.566 ms 0.19% -18033.979 us -76.42% 🟢 FAST
I32 0 2^24 0 16 24.514 ms 7.97% 7.262 ms 0.09% -17251.473 us -70.38% 🟢 FAST
I32 0 2^24 0.1 16 29.784 ms 0.08% 10.797 ms 0.18% -18987.094 us -63.75% 🟢 FAST
I32 0 2^24 0.9 16 10.147 ms 45.33% 8.187 ms 0.28% -1960.384 us -19.32% 🟢 FAST
I32 0 2^24 0 32 16.349 ms 36.36% 11.547 ms 0.44% -4802.145 us -29.37% 🟢 FAST
I32 0 2^24 0.1 32 39.074 ms 0.22% 17.935 ms 0.14% -21139.174 us -54.10% 🟢 FAST
I32 0 2^24 0.9 32 27.693 ms 7.55% 13.383 ms 0.29% -14310.023 us -51.67% 🟢 FAST
F64 0 2^24 0 1 18.115 ms 5.34% 3.428 ms 0.15% -14687.233 us -81.08% 🟢 FAST
F64 0 2^24 0.1 1 24.572 ms 0.08% 3.598 ms 0.15% -20973.979 us -85.36% 🟢 FAST
F64 0 2^24 0.9 1 24.051 ms 0.16% 3.357 ms 0.15% -20693.849 us -86.04% 🟢 FAST
F64 0 2^24 0 2 19.064 ms 5.09% 3.884 ms 0.10% -15179.579 us -79.63% 🟢 FAST
F64 0 2^24 0.1 2 25.534 ms 0.16% 4.198 ms 0.18% -21336.341 us -83.56% 🟢 FAST
F64 0 2^24 0.9 2 8.643 ms 49.44% 3.766 ms 0.15% -4877.531 us -56.43% 🟢 FAST
F64 0 2^24 0 4 8.407 ms 0.20% 4.789 ms 0.09% -3617.717 us -43.03% 🟢 FAST
F64 0 2^24 0.1 4 24.119 ms 8.82% 5.430 ms 0.49% -18688.773 us -77.49% 🟢 FAST
F64 0 2^24 0.9 4 6.462 ms 0.46% 4.601 ms 0.15% -1861.694 us -28.81% 🟢 FAST
F64 0 2^24 0 8 9.161 ms 0.93% 6.609 ms 0.07% -2552.154 us -27.86% 🟢 FAST
F64 0 2^24 0.1 8 10.120 ms 0.17% 7.889 ms 0.22% -2230.806 us -22.04% 🟢 FAST
F64 0 2^24 0.9 8 7.059 ms 0.33% 6.252 ms 0.20% -806.400 us -11.42% 🟢 FAST
F64 0 2^24 0 16 14.859 ms 34.85% 10.229 ms 0.05% -4630.004 us -31.16% 🟢 FAST
F64 0 2^24 0.1 16 16.032 ms 0.23% 12.939 ms 0.11% -3092.814 us -19.29% 🟢 FAST
F64 0 2^24 0.9 16 9.632 ms 0.20% 9.604 ms 0.15% -27.690 us -0.29% 🟢 FAST
F64 0 2^24 0 32 22.225 ms 0.07% 17.456 ms 0.04% -4769.335 us -21.46% 🟢 FAST
F64 0 2^24 0.1 32 20.454 ms 12.78% 22.585 ms 0.06% 2.132 ms 10.42% 🔴 SLOW
F64 0 2^24 0.9 32 18.026 ms 0.08% 16.927 ms 0.19% -1098.664 us -6.09% 🟢 FAST
groupby_struct_keys: 18 faster, 0 slower, 0 same
NumRows Depth Nulls Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
2^10 0 0 115.844 us 4.58% 100.942 us 4.87% -14.902 us -12.86% 🟢 FAST
2^16 0 0 124.781 us 3.39% 105.802 us 4.66% -18.978 us -15.21% 🟢 FAST
2^20 0 0 188.893 us 2.01% 147.544 us 3.09% -41.349 us -21.89% 🟢 FAST
2^10 1 0 260.841 us 2.59% 246.434 us 3.18% -14.407 us -5.52% 🟢 FAST
2^16 1 0 277.128 us 2.87% 257.600 us 3.24% -19.528 us -7.05% 🟢 FAST
2^20 1 0 457.183 us 1.55% 363.042 us 2.10% -94.141 us -20.59% 🟢 FAST
2^10 8 0 983.800 us 1.78% 929.609 us 1.78% -54.191 us -5.51% 🟢 FAST
2^16 8 0 1.007 ms 1.67% 935.984 us 1.57% -70.833 us -7.04% 🟢 FAST
2^20 8 0 1.219 ms 1.38% 1.058 ms 5.17% -160.217 us -13.15% 🟢 FAST
2^10 0 1 158.701 us 3.00% 144.977 us 5.43% -13.725 us -8.65% 🟢 FAST
2^16 0 1 163.521 us 3.31% 150.360 us 3.36% -13.162 us -8.05% 🟢 FAST
2^20 0 1 235.123 us 2.46% 195.138 us 3.86% -39.985 us -17.01% 🟢 FAST
2^10 1 1 259.573 us 2.71% 249.130 us 3.71% -10.443 us -4.02% 🟢 FAST
2^16 1 1 282.175 us 2.81% 255.303 us 2.81% -26.872 us -9.52% 🟢 FAST
2^20 1 1 454.137 us 1.63% 379.451 us 1.69% -74.686 us -16.45% 🟢 FAST
2^10 8 1 986.536 us 1.81% 935.062 us 1.75% -51.474 us -5.22% 🟢 FAST
2^16 8 1 1.017 ms 1.70% 953.532 us 1.64% -63.725 us -6.26% 🟢 FAST
2^20 8 1 1.230 ms 1.46% 1.071 ms 1.64% -158.960 us -12.92% 🟢 FAST
groupby_m2_var_std: 42 faster, 6 slower, 0 same
T U value_key_ratio num_rows null_probability num_aggs Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 12 20 100000 0 1 183.363 us 2.48% 155.387 us 3.42% -27.975 us -15.26% 🟢 FAST
I32 12 100 100000 0 1 171.177 us 2.71% 140.297 us 3.75% -30.880 us -18.04% 🟢 FAST
I32 12 20 10000000 0 1 1.251 ms 0.63% 1.103 ms 0.77% -148.281 us -11.85% 🟢 FAST
I32 12 100 10000000 0 1 1.130 ms 0.53% 659.877 us 0.67% -469.862 us -41.59% 🟢 FAST
I32 12 20 100000 0.5 1 184.286 us 2.75% 160.646 us 3.22% -23.640 us -12.83% 🟢 FAST
I32 12 100 100000 0.5 1 171.166 us 2.83% 140.778 us 3.71% -30.388 us -17.75% 🟢 FAST
I32 12 20 10000000 0.5 1 1.155 ms 0.57% 1.076 ms 0.44% -78.345 us -6.79% 🟢 FAST
I32 12 100 10000000 0.5 1 1.020 ms 0.66% 668.093 us 0.72% -351.653 us -34.48% 🟢 FAST
I32 12 20 100000 0 10 705.481 us 1.62% 644.848 us 1.62% -60.633 us -8.59% 🟢 FAST
I32 12 100 100000 0 10 610.580 us 2.03% 491.392 us 1.82% -119.188 us -19.52% 🟢 FAST
I32 12 20 10000000 0 10 4.799 ms 0.35% 3.743 ms 0.31% -1056.049 us -22.01% 🟢 FAST
I32 12 100 10000000 0 10 4.498 ms 0.18% 2.086 ms 0.37% -2412.348 us -53.63% 🟢 FAST
I32 12 20 100000 0.5 10 742.170 us 1.68% 679.715 us 1.29% -62.455 us -8.42% 🟢 FAST
I32 12 100 100000 0.5 10 646.907 us 1.55% 502.046 us 6.83% -144.861 us -22.39% 🟢 FAST
I32 12 20 10000000 0.5 10 3.808 ms 0.38% 4.129 ms 0.31% 320.368 us 8.41% 🔴 SLOW
I32 12 100 10000000 0.5 10 3.464 ms 0.28% 2.183 ms 0.59% -1280.295 us -36.96% 🟢 FAST
I32 12 20 100000 0 50 2.989 ms 1.01% 2.899 ms 6.78% -89.967 us -3.01% 🟢 FAST
I32 12 100 100000 0 50 2.567 ms 2.10% 2.071 ms 1.32% -495.889 us -19.32% 🟢 FAST
I32 12 20 10000000 0 50 20.493 ms 0.15% 15.367 ms 0.12% -5125.514 us -25.01% 🟢 FAST
I32 12 100 10000000 0 50 19.486 ms 0.17% 8.398 ms 0.24% -11087.416 us -56.90% 🟢 FAST
I32 12 20 100000 0.5 50 3.170 ms 1.43% 2.975 ms 0.99% -195.201 us -6.16% 🟢 FAST
I32 12 100 100000 0.5 50 2.688 ms 1.54% 2.084 ms 2.85% -604.143 us -22.47% 🟢 FAST
I32 12 20 10000000 0.5 50 15.507 ms 0.26% 17.850 ms 0.22% 2.342 ms 15.11% 🔴 SLOW
I32 12 100 10000000 0.5 50 14.263 ms 0.22% 8.874 ms 0.26% -5388.641 us -37.78% 🟢 FAST
F64 12 20 100000 0 1 187.617 us 2.93% 164.240 us 2.83% -23.376 us -12.46% 🟢 FAST
F64 12 100 100000 0 1 177.076 us 3.24% 146.659 us 4.12% -30.417 us -17.18% 🟢 FAST
F64 12 20 10000000 0 1 1.270 ms 0.55% 1.124 ms 0.31% -145.140 us -11.43% 🟢 FAST
F64 12 100 10000000 0 1 1.137 ms 0.73% 862.337 us 0.56% -274.326 us -24.13% 🟢 FAST
F64 12 20 100000 0.5 1 194.707 us 3.46% 166.647 us 2.75% -28.061 us -14.41% 🟢 FAST
F64 12 100 100000 0.5 1 181.488 us 3.03% 145.601 us 3.52% -35.887 us -19.77% 🟢 FAST
F64 12 20 10000000 0.5 1 1.166 ms 0.64% 1.156 ms 0.48% -9.959 us -0.85% 🟢 FAST
F64 12 100 10000000 0.5 1 1.029 ms 0.72% 865.218 us 1.56% -163.831 us -15.92% 🟢 FAST
F64 12 20 100000 0 10 792.794 us 1.94% 735.731 us 9.73% -57.063 us -7.20% 🟢 FAST
F64 12 100 100000 0 10 705.862 us 2.09% 550.068 us 8.64% -155.794 us -22.07% 🟢 FAST
F64 12 20 10000000 0 10 4.953 ms 0.36% 4.598 ms 0.50% -355.113 us -7.17% 🟢 FAST
F64 12 100 10000000 0 10 4.595 ms 0.20% 4.160 ms 0.19% -434.654 us -9.46% 🟢 FAST
F64 12 20 100000 0.5 10 823.053 us 1.80% 754.872 us 1.30% -68.181 us -8.28% 🟢 FAST
F64 12 100 100000 0.5 10 730.991 us 2.26% 537.926 us 2.21% -193.065 us -26.41% 🟢 FAST
F64 12 20 10000000 0.5 10 3.919 ms 0.24% 5.019 ms 0.22% 1.099 ms 28.05% 🔴 SLOW
F64 12 100 10000000 0.5 10 3.568 ms 0.50% 4.150 ms 0.26% 581.549 us 16.30% 🔴 SLOW
F64 12 20 100000 0 50 3.451 ms 2.00% 3.070 ms 1.17% -381.384 us -11.05% 🟢 FAST
F64 12 100 100000 0 50 2.984 ms 1.94% 2.247 ms 1.25% -737.458 us -24.71% 🟢 FAST
F64 12 20 10000000 0 50 21.333 ms 0.25% 19.921 ms 0.08% -1412.628 us -6.62% 🟢 FAST
F64 12 100 10000000 0 50 19.943 ms 0.19% 18.773 ms 0.17% -1170.184 us -5.87% 🟢 FAST
F64 12 20 100000 0.5 50 3.632 ms 1.46% 3.332 ms 0.78% -300.128 us -8.26% 🟢 FAST
F64 12 100 100000 0.5 50 3.122 ms 1.00% 2.267 ms 1.30% -855.265 us -27.40% 🟢 FAST
F64 12 20 10000000 0.5 50 16.086 ms 0.30% 21.946 ms 0.13% 5.860 ms 36.43% 🔴 SLOW
F64 12 100 10000000 0.5 50 14.747 ms 0.22% 18.755 ms 0.12% 4.008 ms 27.18% 🔴 SLOW
groupby_m2_var_std, extra axis: 32 to 200 rows per group: 13 faster, 7 slower, 0 same
T U value_key_ratio num_rows null_probability num_aggs Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 12 32 10000000 0 10 4.575 ms 0.25% 3.410 ms 0.28% -1165.055 us -25.47% 🟢 FAST
I32 12 50 10000000 0 10 4.521 ms 0.25% 3.322 ms 0.31% -1198.886 us -26.52% 🟢 FAST
I32 12 64 10000000 0 10 4.523 ms 0.26% 3.281 ms 0.29% -1242.304 us -27.47% 🟢 FAST
I32 12 100 10000000 0 10 4.505 ms 0.22% 2.095 ms 0.41% -2410.290 us -53.50% 🟢 FAST
I32 12 200 10000000 0 10 4.578 ms 0.30% 1.638 ms 0.45% -2940.638 us -64.23% 🟢 FAST
I32 12 32 10000000 0.5 10 3.586 ms 0.23% 3.868 ms 0.24% 282.178 us 7.87% 🔴 SLOW
I32 12 50 10000000 0.5 10 3.519 ms 0.50% 3.777 ms 0.19% 257.415 us 7.31% 🔴 SLOW
I32 12 64 10000000 0.5 10 3.496 ms 0.24% 3.729 ms 0.20% 232.809 us 6.66% 🔴 SLOW
I32 12 100 10000000 0.5 10 3.478 ms 0.34% 2.180 ms 0.42% -1298.252 us -37.33% 🟢 FAST
I32 12 200 10000000 0.5 10 3.542 ms 1.02% 1.664 ms 0.42% -1878.482 us -53.03% 🟢 FAST
F64 12 32 10000000 0 10 4.703 ms 0.19% 4.259 ms 0.25% -443.920 us -9.44% 🟢 FAST
F64 12 50 10000000 0 10 4.649 ms 0.50% 4.169 ms 0.50% -479.602 us -10.32% 🟢 FAST
F64 12 64 10000000 0 10 4.648 ms 0.24% 4.126 ms 0.24% -521.764 us -11.23% 🟢 FAST
F64 12 100 10000000 0 10 4.618 ms 0.19% 4.158 ms 0.22% -460.467 us -9.97% 🟢 FAST
F64 12 200 10000000 0 10 4.710 ms 0.19% 2.774 ms 0.31% -1936.038 us -41.11% 🟢 FAST
F64 12 32 10000000 0.5 10 3.690 ms 0.36% 4.666 ms 0.23% 975.751 us 26.44% 🔴 SLOW
F64 12 50 10000000 0.5 10 3.622 ms 0.26% 4.584 ms 0.24% 961.694 us 26.55% 🔴 SLOW
F64 12 64 10000000 0.5 10 3.613 ms 0.31% 4.557 ms 0.17% 943.781 us 26.12% 🔴 SLOW
F64 12 100 10000000 0.5 10 3.594 ms 0.28% 4.154 ms 0.20% 559.553 us 15.57% 🔴 SLOW
F64 12 200 10000000 0.5 10 3.607 ms 0.40% 2.755 ms 0.38% -852.190 us -23.63% 🟢 FAST
sum: 6 faster, 2 slower, 0 same
T num_rows Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I64 100000 136.001 us 3.10% 113.854 us 3.90% -22.147 us -16.28% 🟢 FAST
I64 1000000 203.104 us 2.09% 136.065 us 4.53% -67.038 us -33.01% 🟢 FAST
I64 10000000 793.297 us 1.06% 562.268 us 0.64% -231.029 us -29.12% 🟢 FAST
I64 100000000 5.539 ms 0.31% 5.897 ms 0.16% 357.312 us 6.45% 🔴 SLOW
decimal64 100000 140.317 us 4.59% 116.888 us 3.86% -23.428 us -16.70% 🟢 FAST
decimal64 1000000 352.257 us 2.99% 167.914 us 2.31% -184.342 us -52.33% 🟢 FAST
decimal64 10000000 2.381 ms 0.50% 1.713 ms 0.19% -667.998 us -28.06% 🟢 FAST
decimal64 100000000 26.328 ms 0.15% 30.495 ms 0.07% 4.167 ms 15.83% 🔴 SLOW
no_requests: 3 faster, 0 slower, 1 same
num_rows Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
100000 47.422 us 7.48% 45.536 us 6.74% -1.887 us -3.98% 🔵 SAME
1000000 93.446 us 2.72% 56.542 us 4.89% -36.904 us -39.49% 🟢 FAST
10000000 340.539 us 0.90% 211.472 us 1.43% -129.068 us -37.90% 🟢 FAST
100000000 2.842 ms 0.28% 1.835 ms 0.22% -1007.541 us -35.45% 🟢 FAST

@PointKernel

PointKernel commented Sep 9, 2026

Copy link
Copy Markdown
Member Author

Groupby benchmarks vs main (2062f4d) at 11d3dd1, RTX PRO 6000 (sm_120), CUDA 12.9, idle GPU, nvbench_compare.py status relative to noise. Supersedes the earlier comment, whose main times for groupby_max, complex_int_keys and the 20 to 1M group configs were inflated by another job on the GPU.

benchmark faster slower same
groupby_max_cardinality, 20 to 1M groups 54 1 1
groupby_max_cardinality, 1 to 16 groups 18 2 0
groupby_max_cardinality, 2M to 5M groups 8 0 0
complex_int_keys 20 0 0
complex_mixed_keys 33 7 0
groupby_max 31 5 0
groupby_struct_keys 18 0 0
groupby_m2_var_std, default axes 48 0 0
groupby_m2_var_std, 32 to 1000 rows per group 56 0 0
sum 7 1 0
no_requests 3 0 1
total 296 16 2

Since 67c719a: hot table slots are read through L1; for 2M+ rows with keys up to 32 bytes the table is sized from a sampled distinct-key estimate (full-size rebuild if it falls short); groups averaging under 128 rows are packed several per block via the segmented-reduce size hint, replacing ReduceByKey and its label array; nullable aggregations produce the result and its null mask in one pass.

case (ms) main previous HashCSR (67c719a) now peak MB, previous HashCSR -> now
20M rows, 1 group, 1 MAX 1.053 1.033 0.682 560 -> 321
20M rows, 8 groups, 8 MAX 1.729 2.250 1.900 560 -> 321
20M rows, 100 groups, 1 MAX 1.263 1.047 0.797 560 -> 321
20M rows, 1M groups, 8 MAX 4.783 5.112 3.043 564 -> 339
VAR/STD 10M F64, 20 rows/group, 50% nulls, 10 aggs 3.920 4.978 2.597 282 -> 213
groupby_max unique keys, 90% nulls, 32 aggs 7.129 13.329 7.652 575 -> 519
no_requests 100M rows 2.826 1.826 0.918 1200 -> 401

Slower than main: 8 to 20 groups with 8 aggregations (6 to 10%), groupby_max with 90% null keys and values at 16 to 32 aggregations (up to 28%), sum decimal64 at 100M near-unique keys (13%), 3 to 5 mixed int/string/list key columns (3 to 8%). Peak memory is 1.7x below main on the 20M-row benchmarks, 3x for 100M-row distinct keys, equal for unique and wide keys.

groupby_max_cardinality, 20 to 1M groups: 54 faster, 1 slower, 1 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 20 normal 1.031 ms 0.66% 779.602 us 0.80% -251.135 us -24.36% 🟢 FAST
I32 20000000 2 20 normal 1.150 ms 2.04% 955.986 us 0.79% -194.118 us -16.88% 🟢 FAST
I32 20000000 4 20 normal 1.400 ms 1.38% 1.274 ms 0.55% -125.893 us -8.99% 🟢 FAST
I32 20000000 8 20 normal 1.798 ms 1.46% 1.918 ms 0.50% 120.078 us 6.68% 🔴 SLOW
I32 20000000 1 50 normal 1.118 ms 0.68% 785.541 us 0.74% -332.715 us -29.75% 🟢 FAST
I32 20000000 2 50 normal 1.274 ms 0.64% 960.716 us 0.52% -313.507 us -24.60% 🟢 FAST
I32 20000000 4 50 normal 1.532 ms 0.93% 1.304 ms 0.46% -228.112 us -14.89% 🟢 FAST
I32 20000000 8 50 normal 1.969 ms 1.10% 1.971 ms 0.39% 1.985 us 0.10% 🔵 SAME
I32 20000000 1 100 normal 1.263 ms 0.50% 796.551 us 0.67% -466.026 us -36.91% 🟢 FAST
I32 20000000 2 100 normal 1.425 ms 0.62% 969.108 us 0.52% -455.997 us -32.00% 🟢 FAST
I32 20000000 4 100 normal 1.708 ms 1.07% 1.325 ms 0.45% -383.385 us -22.44% 🟢 FAST
I32 20000000 8 100 normal 2.152 ms 0.88% 2.006 ms 0.31% -145.652 us -6.77% 🟢 FAST
I32 20000000 1 1000 normal 2.785 ms 2.24% 801.073 us 0.46% -1983.725 us -71.23% 🟢 FAST
I32 20000000 2 1000 normal 3.178 ms 4.05% 976.514 us 0.41% -2201.731 us -69.28% 🟢 FAST
I32 20000000 4 1000 normal 7.027 ms 0.22% 1.313 ms 0.36% -5714.319 us -81.31% 🟢 FAST
I32 20000000 8 1000 normal 11.665 ms 0.12% 1.981 ms 0.27% -9684.763 us -83.02% 🟢 FAST
I32 20000000 1 10000 normal 2.701 ms 0.77% 845.381 us 0.56% -1855.346 us -68.70% 🟢 FAST
I32 20000000 2 10000 normal 2.928 ms 0.80% 1.021 ms 0.35% -1906.260 us -65.11% 🟢 FAST
I32 20000000 4 10000 normal 3.679 ms 0.44% 1.346 ms 0.29% -2332.672 us -63.40% 🟢 FAST
I32 20000000 8 10000 normal 5.132 ms 0.20% 2.000 ms 0.21% -3131.985 us -61.02% 🟢 FAST
I32 20000000 1 100000 normal 2.730 ms 0.27% 1.090 ms 0.41% -1640.612 us -60.09% 🟢 FAST
I32 20000000 2 100000 normal 2.963 ms 0.26% 1.296 ms 0.35% -1666.912 us -56.26% 🟢 FAST
I32 20000000 4 100000 normal 3.265 ms 0.19% 1.695 ms 0.23% -1569.639 us -48.07% 🟢 FAST
I32 20000000 8 100000 normal 4.151 ms 0.22% 2.487 ms 0.16% -1663.882 us -40.09% 🟢 FAST
I32 20000000 1 1000000 normal 3.991 ms 0.32% 1.470 ms 0.50% -2521.151 us -63.17% 🟢 FAST
I32 20000000 2 1000000 normal 5.433 ms 0.50% 1.682 ms 0.25% -3751.341 us -69.05% 🟢 FAST
I32 20000000 4 1000000 normal 3.871 ms 0.16% 2.085 ms 0.23% -1785.625 us -46.13% 🟢 FAST
I32 20000000 8 1000000 normal 4.783 ms 0.50% 3.043 ms 0.29% -1739.311 us -36.37% 🟢 FAST
decimal128 20000000 1 20 normal 82.847 ms 0.19% 961.751 us 0.46% -81884.874 us -98.84% 🟢 FAST
decimal128 20000000 2 20 normal 83.951 ms 0.14% 1.342 ms 0.40% -82609.389 us -98.40% 🟢 FAST
decimal128 20000000 4 20 normal 85.020 ms 0.14% 2.085 ms 0.36% -82934.811 us -97.55% 🟢 FAST
decimal128 20000000 8 20 normal 87.788 ms 0.23% 3.568 ms 0.24% -84220.149 us -95.94% 🟢 FAST
decimal128 20000000 1 50 normal 85.050 ms 0.16% 974.236 us 0.55% -84076.147 us -98.85% 🟢 FAST
decimal128 20000000 2 50 normal 85.939 ms 0.16% 1.366 ms 0.47% -84572.696 us -98.41% 🟢 FAST
decimal128 20000000 4 50 normal 87.759 ms 0.24% 2.137 ms 0.30% -85622.365 us -97.57% 🟢 FAST
decimal128 20000000 8 50 normal 91.472 ms 0.18% 3.679 ms 0.28% -87793.150 us -95.98% 🟢 FAST
decimal128 20000000 1 100 normal 87.719 ms 0.16% 1.002 ms 0.60% -86716.389 us -98.86% 🟢 FAST
decimal128 20000000 2 100 normal 88.648 ms 0.17% 1.400 ms 0.38% -87248.214 us -98.42% 🟢 FAST
decimal128 20000000 4 100 normal 90.471 ms 0.29% 2.177 ms 0.28% -88293.286 us -97.59% 🟢 FAST
decimal128 20000000 8 100 normal 94.270 ms 0.10% 3.736 ms 0.18% -90534.373 us -96.04% 🟢 FAST
decimal128 20000000 1 1000 normal 94.525 ms 0.20% 1.028 ms 0.47% -93496.667 us -98.91% 🟢 FAST
decimal128 20000000 2 1000 normal 95.248 ms 0.13% 1.424 ms 0.36% -93824.578 us -98.51% 🟢 FAST
decimal128 20000000 4 1000 normal 97.034 ms 0.13% 2.200 ms 0.24% -94833.505 us -97.73% 🟢 FAST
decimal128 20000000 8 1000 normal 100.969 ms 0.11% 3.762 ms 0.26% -97206.930 us -96.27% 🟢 FAST
decimal128 20000000 1 10000 normal 96.014 ms 0.15% 1.068 ms 0.50% -94946.006 us -98.89% 🟢 FAST
decimal128 20000000 2 10000 normal 96.848 ms 0.13% 1.451 ms 0.32% -95396.862 us -98.50% 🟢 FAST
decimal128 20000000 4 10000 normal 98.707 ms 0.10% 2.210 ms 0.25% -96496.171 us -97.76% 🟢 FAST
decimal128 20000000 8 10000 normal 102.552 ms 0.09% 3.723 ms 0.17% -98829.071 us -96.37% 🟢 FAST
decimal128 20000000 1 100000 normal 89.228 ms 0.14% 1.260 ms 0.68% -87968.247 us -98.59% 🟢 FAST
decimal128 20000000 2 100000 normal 90.163 ms 0.14% 1.607 ms 0.33% -88555.971 us -98.22% 🟢 FAST
decimal128 20000000 4 100000 normal 92.171 ms 0.12% 2.307 ms 0.34% -89864.681 us -97.50% 🟢 FAST
decimal128 20000000 8 100000 normal 96.197 ms 0.12% 3.706 ms 0.29% -92491.419 us -96.15% 🟢 FAST
decimal128 20000000 1 1000000 normal 79.836 ms 0.16% 1.748 ms 0.30% -78088.364 us -97.81% 🟢 FAST
decimal128 20000000 2 1000000 normal 80.778 ms 0.15% 2.206 ms 0.21% -78571.430 us -97.27% 🟢 FAST
decimal128 20000000 4 1000000 normal 82.510 ms 0.12% 3.118 ms 0.20% -79391.902 us -96.22% 🟢 FAST
decimal128 20000000 8 1000000 normal 86.134 ms 0.18% 4.946 ms 0.20% -81187.663 us -94.26% 🟢 FAST
groupby_max_cardinality, 1 to 16 groups: 18 faster, 2 slower, 0 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 1 normal 1.053 ms 2.01% 681.977 us 1.17% -370.906 us -35.23% 🟢 FAST
I32 20000000 8 1 normal 1.663 ms 0.83% 1.387 ms 0.54% -276.584 us -16.63% 🟢 FAST
I32 20000000 1 2 normal 1.036 ms 0.64% 704.900 us 0.83% -331.244 us -31.97% 🟢 FAST
I32 20000000 8 2 normal 1.697 ms 1.21% 1.502 ms 0.44% -194.581 us -11.47% 🟢 FAST
I32 20000000 1 4 normal 1.038 ms 0.63% 730.812 us 0.58% -306.702 us -29.56% 🟢 FAST
I32 20000000 8 4 normal 1.674 ms 1.58% 1.624 ms 0.40% -49.968 us -2.99% 🟢 FAST
I32 20000000 1 8 normal 1.039 ms 0.59% 752.338 us 0.66% -286.844 us -27.60% 🟢 FAST
I32 20000000 8 8 normal 1.729 ms 1.83% 1.900 ms 0.50% 170.509 us 9.86% 🔴 SLOW
I32 20000000 1 16 normal 1.040 ms 0.75% 787.607 us 0.80% -252.517 us -24.28% 🟢 FAST
I32 20000000 8 16 normal 1.789 ms 1.48% 1.966 ms 0.33% 176.855 us 9.89% 🔴 SLOW
decimal128 20000000 1 1 normal 55.772 ms 0.20% 844.116 us 1.01% -54928.384 us -98.49% 🟢 FAST
decimal128 20000000 8 1 normal 61.195 ms 0.13% 2.629 ms 0.38% -58566.126 us -95.70% 🟢 FAST
decimal128 20000000 1 2 normal 67.851 ms 0.26% 870.708 us 0.62% -66980.316 us -98.72% 🟢 FAST
decimal128 20000000 8 2 normal 73.712 ms 0.17% 2.759 ms 0.27% -70952.491 us -96.26% 🟢 FAST
decimal128 20000000 1 4 normal 73.118 ms 0.17% 895.942 us 0.65% -72221.898 us -98.77% 🟢 FAST
decimal128 20000000 8 4 normal 79.045 ms 0.20% 3.006 ms 0.22% -76039.270 us -96.20% 🟢 FAST
decimal128 20000000 1 8 normal 78.290 ms 0.23% 936.371 us 1.36% -77353.527 us -98.80% 🟢 FAST
decimal128 20000000 8 8 normal 84.458 ms 0.22% 3.327 ms 0.50% -81131.019 us -96.06% 🟢 FAST
decimal128 20000000 1 16 normal 80.845 ms 0.21% 961.453 us 0.57% -79883.394 us -98.81% 🟢 FAST
decimal128 20000000 8 16 normal 87.248 ms 0.21% 3.532 ms 0.22% -83716.084 us -95.95% 🟢 FAST
groupby_max_cardinality, 2M to 5M groups: 8 faster, 0 slower, 0 same
T num_rows num_aggregations cardinality api Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 20000000 1 2000000 normal 4.327 ms 0.62% 1.684 ms 2.12% -2642.909 us -61.08% 🟢 FAST
I32 20000000 8 2000000 normal 5.214 ms 0.20% 3.109 ms 0.28% -2105.287 us -40.38% 🟢 FAST
I32 20000000 1 5000000 normal 4.480 ms 0.29% 3.613 ms 0.50% -866.469 us -19.34% 🟢 FAST
I32 20000000 8 5000000 normal 5.757 ms 0.16% 5.068 ms 0.12% -688.981 us -11.97% 🟢 FAST
decimal128 20000000 1 2000000 normal 76.067 ms 0.12% 1.833 ms 0.25% -74233.576 us -97.59% 🟢 FAST
decimal128 20000000 8 2000000 normal 83.085 ms 0.17% 4.877 ms 0.13% -78208.109 us -94.13% 🟢 FAST
decimal128 20000000 1 5000000 normal 71.242 ms 0.11% 3.819 ms 0.15% -67423.531 us -94.64% 🟢 FAST
decimal128 20000000 8 5000000 normal 80.014 ms 0.10% 6.661 ms 0.50% -73352.895 us -91.68% 🟢 FAST
complex_int_keys: 20 faster, 0 slower, 0 same
num_cols num_rows value_key_ratio null_probability Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
1 2^24 20 0 3.340 ms 0.79% 1.271 ms 3.64% -2068.754 us -61.95% 🟢 FAST
2 2^24 20 0 3.735 ms 0.24% 1.428 ms 0.73% -2306.632 us -61.76% 🟢 FAST
4 2^24 20 0 4.422 ms 0.25% 1.823 ms 2.28% -2599.293 us -58.78% 🟢 FAST
8 2^24 20 0 6.194 ms 0.23% 2.929 ms 0.38% -3265.026 us -52.72% 🟢 FAST
16 2^24 20 0 11.311 ms 0.19% 5.328 ms 0.07% -5982.528 us -52.89% 🟢 FAST
1 2^24 200 0 1.864 ms 0.45% 845.591 us 0.42% -1017.966 us -54.62% 🟢 FAST
2 2^24 200 0 2.098 ms 0.45% 982.134 us 0.55% -1116.355 us -53.20% 🟢 FAST
4 2^24 200 0 2.797 ms 0.50% 1.262 ms 0.90% -1534.550 us -54.87% 🟢 FAST
8 2^24 200 0 4.981 ms 0.23% 1.807 ms 0.28% -3173.746 us -63.72% 🟢 FAST
16 2^24 200 0 8.991 ms 0.19% 2.452 ms 0.18% -6539.407 us -72.73% 🟢 FAST
1 2^24 20 0.5 2.297 ms 0.50% 787.183 us 0.77% -1509.751 us -65.73% 🟢 FAST
2 2^24 20 0.5 2.647 ms 0.37% 1.041 ms 1.95% -1605.547 us -60.66% 🟢 FAST
4 2^24 20 0.5 3.168 ms 0.48% 1.475 ms 0.59% -1693.258 us -53.44% 🟢 FAST
8 2^24 20 0.5 5.105 ms 0.24% 2.434 ms 0.34% -2670.819 us -52.32% 🟢 FAST
16 2^24 20 0.5 9.058 ms 0.50% 4.136 ms 0.73% -4922.456 us -54.34% 🟢 FAST
1 2^24 200 0.5 1.787 ms 0.34% 668.577 us 2.94% -1118.329 us -62.58% 🟢 FAST
2 2^24 200 0.5 2.047 ms 0.43% 883.099 us 2.37% -1164.214 us -56.87% 🟢 FAST
4 2^24 200 0.5 2.665 ms 0.28% 1.237 ms 0.63% -1427.572 us -53.57% 🟢 FAST
8 2^24 200 0.5 4.763 ms 0.64% 1.969 ms 0.50% -2793.983 us -58.67% 🟢 FAST
16 2^24 200 0.5 8.392 ms 0.30% 3.116 ms 0.25% -5275.838 us -62.87% 🟢 FAST
complex_mixed_keys: 33 faster, 7 slower, 0 same
num_cols num_rows value_key_ratio null_probability Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
1 2^18 20 0 154.960 us 10.20% 126.675 us 9.93% -28.285 us -18.25% 🟢 FAST
2 2^18 20 0 241.995 us 5.98% 210.609 us 5.34% -31.386 us -12.97% 🟢 FAST
3 2^18 20 0 306.846 us 2.71% 230.753 us 8.75% -76.092 us -24.80% 🟢 FAST
4 2^18 20 0 3.130 ms 0.25% 2.838 ms 0.96% -291.860 us -9.33% 🟢 FAST
5 2^18 20 0 3.293 ms 0.30% 2.947 ms 0.33% -346.748 us -10.53% 🟢 FAST
1 2^24 20 0 3.305 ms 0.32% 1.263 ms 0.46% -2041.437 us -61.77% 🟢 FAST
2 2^24 20 0 7.237 ms 1.46% 6.942 ms 0.98% -294.671 us -4.07% 🟢 FAST
3 2^24 20 0 8.622 ms 0.21% 8.986 ms 0.12% 363.891 us 4.22% 🔴 SLOW
4 2^24 20 0 42.136 ms 0.09% 44.717 ms 0.10% 2.582 ms 6.13% 🔴 SLOW
5 2^24 20 0 48.216 ms 0.32% 49.503 ms 0.14% 1.287 ms 2.67% 🔴 SLOW
1 2^18 200 0 142.923 us 8.11% 112.304 us 3.63% -30.619 us -21.42% 🟢 FAST
2 2^18 200 0 244.399 us 2.29% 207.918 us 2.72% -36.481 us -14.93% 🟢 FAST
3 2^18 200 0 303.164 us 2.32% 224.476 us 2.23% -78.689 us -25.96% 🟢 FAST
4 2^18 200 0 3.135 ms 0.26% 2.834 ms 0.30% -301.279 us -9.61% 🟢 FAST
5 2^18 200 0 3.297 ms 0.34% 2.946 ms 0.35% -350.655 us -10.64% 🟢 FAST
1 2^24 200 0 1.850 ms 0.51% 839.007 us 0.43% -1011.319 us -54.66% 🟢 FAST
2 2^24 200 0 6.689 ms 0.50% 6.778 ms 0.50% 88.615 us 1.32% 🔴 SLOW
3 2^24 200 0 8.352 ms 0.18% 9.026 ms 0.38% 674.519 us 8.08% 🔴 SLOW
4 2^24 200 0 42.063 ms 0.12% 44.750 ms 0.10% 2.687 ms 6.39% 🔴 SLOW
5 2^24 200 0 48.085 ms 0.22% 49.481 ms 0.11% 1.396 ms 2.90% 🔴 SLOW
1 2^18 20 0.5 222.503 us 3.02% 191.385 us 2.77% -31.118 us -13.99% 🟢 FAST
2 2^18 20 0.5 355.784 us 2.70% 319.730 us 4.32% -36.054 us -10.13% 🟢 FAST
3 2^18 20 0.5 356.929 us 2.90% 326.428 us 2.69% -30.501 us -8.55% 🟢 FAST
4 2^18 20 0.5 643.305 us 1.16% 560.553 us 1.74% -82.752 us -12.86% 🟢 FAST
5 2^18 20 0.5 1.201 ms 0.83% 1.101 ms 3.54% -99.766 us -8.31% 🟢 FAST
1 2^24 20 0.5 2.346 ms 0.39% 805.066 us 1.16% -1540.879 us -65.68% 🟢 FAST
2 2^24 20 0.5 3.244 ms 0.44% 2.236 ms 0.53% -1008.334 us -31.08% 🟢 FAST
3 2^24 20 0.5 2.631 ms 0.48% 1.876 ms 0.42% -754.268 us -28.67% 🟢 FAST
4 2^24 20 0.5 3.082 ms 0.43% 2.827 ms 0.29% -255.676 us -8.29% 🟢 FAST
5 2^24 20 0.5 3.458 ms 0.35% 3.209 ms 0.36% -248.629 us -7.19% 🟢 FAST
1 2^18 200 0.5 220.679 us 6.70% 179.353 us 4.78% -41.326 us -18.73% 🟢 FAST
2 2^18 200 0.5 354.712 us 1.55% 323.388 us 6.50% -31.324 us -8.83% 🟢 FAST
3 2^18 200 0.5 359.169 us 5.07% 320.675 us 2.13% -38.495 us -10.72% 🟢 FAST
4 2^18 200 0.5 648.821 us 1.07% 560.118 us 1.36% -88.702 us -13.67% 🟢 FAST
5 2^18 200 0.5 1.227 ms 4.48% 1.102 ms 1.04% -124.767 us -10.17% 🟢 FAST
1 2^24 200 0.5 1.805 ms 0.42% 675.942 us 0.85% -1128.796 us -62.55% 🟢 FAST
2 2^24 200 0.5 3.248 ms 0.68% 2.265 ms 1.98% -982.657 us -30.25% 🟢 FAST
3 2^24 200 0.5 2.626 ms 0.40% 1.882 ms 1.84% -743.581 us -28.32% 🟢 FAST
4 2^24 200 0.5 3.078 ms 0.33% 2.842 ms 1.31% -236.693 us -7.69% 🟢 FAST
5 2^24 200 0.5 3.452 ms 0.37% 3.249 ms 2.37% -203.313 us -5.89% 🟢 FAST
groupby_max: 31 faster, 5 slower, 0 same
T cardinality num_rows null_probability num_aggregations Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 0 2^24 0 1 3.717 ms 1.63% 3.039 ms 0.50% -678.355 us -18.25% 🟢 FAST
I32 0 2^24 0.1 1 3.796 ms 0.26% 3.098 ms 0.55% -697.917 us -18.39% 🟢 FAST
I32 0 2^24 0.9 1 3.423 ms 0.36% 2.995 ms 0.34% -428.126 us -12.51% 🟢 FAST
I32 0 2^24 0 2 4.376 ms 0.27% 3.200 ms 0.13% -1175.587 us -26.86% 🟢 FAST
I32 0 2^24 0.1 2 4.539 ms 0.32% 3.350 ms 0.14% -1189.045 us -26.20% 🟢 FAST
I32 0 2^24 0.9 2 3.789 ms 0.31% 3.143 ms 0.23% -645.685 us -17.04% 🟢 FAST
I32 0 2^24 0 4 4.049 ms 0.29% 3.537 ms 0.14% -512.034 us -12.65% 🟢 FAST
I32 0 2^24 0.1 4 4.468 ms 0.25% 3.853 ms 0.16% -614.693 us -13.76% 🟢 FAST
I32 0 2^24 0.9 4 3.769 ms 0.48% 3.446 ms 0.26% -323.402 us -8.58% 🟢 FAST
I32 0 2^24 0 8 4.860 ms 0.22% 4.221 ms 0.13% -638.119 us -13.13% 🟢 FAST
I32 0 2^24 0.1 8 5.698 ms 0.40% 4.871 ms 0.50% -826.607 us -14.51% 🟢 FAST
I32 0 2^24 0.9 8 4.251 ms 0.50% 4.049 ms 0.17% -201.955 us -4.75% 🟢 FAST
I32 0 2^24 0 16 6.533 ms 0.21% 5.597 ms 0.10% -936.321 us -14.33% 🟢 FAST
I32 0 2^24 0.1 16 8.209 ms 0.22% 6.905 ms 0.26% -1303.738 us -15.88% 🟢 FAST
I32 0 2^24 0.9 16 5.240 ms 0.28% 5.252 ms 0.17% 11.684 us 0.22% 🔴 SLOW
I32 0 2^24 0 32 9.830 ms 0.22% 8.327 ms 0.06% -1503.168 us -15.29% 🟢 FAST
I32 0 2^24 0.1 32 13.158 ms 0.21% 10.957 ms 0.12% -2201.136 us -16.73% 🟢 FAST
I32 0 2^24 0.9 32 7.129 ms 0.16% 7.652 ms 0.16% 522.662 us 7.33% 🔴 SLOW
F64 0 2^24 0 1 3.983 ms 0.25% 3.204 ms 0.50% -778.730 us -19.55% 🟢 FAST
F64 0 2^24 0.1 1 4.099 ms 0.24% 3.160 ms 0.17% -938.696 us -22.90% 🟢 FAST
F64 0 2^24 0.9 1 3.544 ms 0.28% 3.054 ms 0.16% -490.820 us -13.85% 🟢 FAST
F64 0 2^24 0 2 4.928 ms 0.26% 3.525 ms 0.18% -1403.742 us -28.48% 🟢 FAST
F64 0 2^24 0.1 2 5.067 ms 0.29% 3.455 ms 0.53% -1612.264 us -31.82% 🟢 FAST
F64 0 2^24 0.9 2 4.007 ms 0.20% 3.280 ms 0.16% -727.242 us -18.15% 🟢 FAST
F64 0 2^24 0 4 4.889 ms 0.23% 4.179 ms 0.15% -710.446 us -14.53% 🟢 FAST
F64 0 2^24 0.1 4 5.219 ms 0.24% 4.047 ms 0.25% -1171.470 us -22.45% 🟢 FAST
F64 0 2^24 0.9 4 3.896 ms 0.32% 3.719 ms 0.16% -176.901 us -4.54% 🟢 FAST
F64 0 2^24 0 8 6.584 ms 0.16% 5.475 ms 0.50% -1109.569 us -16.85% 🟢 FAST
F64 0 2^24 0.1 8 7.200 ms 0.23% 5.236 ms 0.13% -1963.673 us -27.27% 🟢 FAST
F64 0 2^24 0.9 8 4.515 ms 0.32% 4.642 ms 0.46% 126.694 us 2.81% 🔴 SLOW
F64 0 2^24 0 16 9.982 ms 0.27% 8.065 ms 0.11% -1917.493 us -19.21% 🟢 FAST
F64 0 2^24 0.1 16 11.206 ms 0.35% 7.601 ms 0.13% -3604.794 us -32.17% 🟢 FAST
F64 0 2^24 0.9 16 5.738 ms 0.23% 6.444 ms 0.15% 705.902 us 12.30% 🔴 SLOW
F64 0 2^24 0 32 16.862 ms 0.50% 13.261 ms 0.15% -3601.064 us -21.36% 🟢 FAST
F64 0 2^24 0.1 32 19.403 ms 0.50% 12.367 ms 0.13% -7036.192 us -36.26% 🟢 FAST
F64 0 2^24 0.9 32 8.217 ms 0.50% 10.546 ms 0.30% 2.329 ms 28.34% 🔴 SLOW
groupby_struct_keys: 18 faster, 0 slower, 0 same
NumRows Depth Nulls Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
2^10 0 0 121.149 us 10.25% 105.327 us 6.52% -15.822 us -13.06% 🟢 FAST
2^16 0 0 131.008 us 5.83% 107.466 us 4.01% -23.542 us -17.97% 🟢 FAST
2^20 0 0 189.628 us 2.65% 142.875 us 2.61% -46.753 us -24.66% 🟢 FAST
2^10 1 0 261.206 us 2.33% 249.380 us 2.40% -11.827 us -4.53% 🟢 FAST
2^16 1 0 281.101 us 2.51% 257.570 us 2.63% -23.531 us -8.37% 🟢 FAST
2^20 1 0 456.862 us 3.24% 370.005 us 1.74% -86.857 us -19.01% 🟢 FAST
2^10 8 0 997.070 us 1.11% 951.140 us 0.84% -45.930 us -4.61% 🟢 FAST
2^16 8 0 1.008 ms 1.26% 963.634 us 1.14% -44.000 us -4.37% 🟢 FAST
2^20 8 0 1.227 ms 1.06% 1.078 ms 0.94% -149.286 us -12.17% 🟢 FAST
2^10 0 1 159.324 us 3.28% 149.436 us 3.43% -9.888 us -6.21% 🟢 FAST
2^16 0 1 164.664 us 2.60% 151.402 us 3.04% -13.262 us -8.05% 🟢 FAST
2^20 0 1 233.623 us 5.12% 194.250 us 2.48% -39.373 us -16.85% 🟢 FAST
2^10 1 1 260.933 us 2.39% 249.189 us 2.60% -11.744 us -4.50% 🟢 FAST
2^16 1 1 284.515 us 4.25% 258.582 us 3.18% -25.934 us -9.12% 🟢 FAST
2^20 1 1 457.147 us 2.74% 374.030 us 1.48% -83.116 us -18.18% 🟢 FAST
2^10 8 1 994.232 us 1.07% 945.874 us 2.30% -48.357 us -4.86% 🟢 FAST
2^16 8 1 1.018 ms 1.84% 964.457 us 1.18% -54.009 us -5.30% 🟢 FAST
2^20 8 1 1.230 ms 1.16% 1.074 ms 0.86% -155.961 us -12.68% 🟢 FAST
groupby_m2_var_std, default axes: 48 faster, 0 slower, 0 same
T U value_key_ratio num_rows null_probability num_aggs Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 12 20 100000 0 1 186.794 us 6.27% 160.707 us 9.45% -26.086 us -13.97% 🟢 FAST
I32 12 100 100000 0 1 174.280 us 3.65% 154.582 us 6.57% -19.698 us -11.30% 🟢 FAST
I32 12 20 10000000 0 1 1.250 ms 0.64% 626.818 us 3.07% -623.454 us -49.87% 🟢 FAST
I32 12 100 10000000 0 1 1.124 ms 0.62% 499.454 us 2.80% -625.028 us -55.58% 🟢 FAST
I32 12 20 100000 0.5 1 187.310 us 2.64% 161.921 us 2.93% -25.390 us -13.55% 🟢 FAST
I32 12 100 100000 0.5 1 175.431 us 3.35% 160.142 us 2.56% -15.290 us -8.72% 🟢 FAST
I32 12 20 10000000 0.5 1 1.154 ms 0.60% 624.458 us 0.74% -529.153 us -45.87% 🟢 FAST
I32 12 100 10000000 0.5 1 1.016 ms 0.70% 505.223 us 0.90% -510.565 us -50.26% 🟢 FAST
I32 12 20 100000 0 10 714.778 us 4.04% 594.791 us 1.18% -119.988 us -16.79% 🟢 FAST
I32 12 100 100000 0 10 630.941 us 4.16% 558.821 us 1.31% -72.120 us -11.43% 🟢 FAST
I32 12 20 10000000 0 10 4.788 ms 0.50% 1.895 ms 0.92% -2893.354 us -60.43% 🟢 FAST
I32 12 100 10000000 0 10 4.478 ms 0.48% 1.430 ms 1.68% -3047.924 us -68.06% 🟢 FAST
I32 12 20 100000 0.5 10 777.527 us 8.97% 620.617 us 0.97% -156.909 us -20.18% 🟢 FAST
I32 12 100 100000 0.5 10 701.655 us 11.61% 601.678 us 1.93% -99.976 us -14.25% 🟢 FAST
I32 12 20 10000000 0.5 10 3.788 ms 0.36% 1.957 ms 0.50% -1831.277 us -48.34% 🟢 FAST
I32 12 100 10000000 0.5 10 3.448 ms 0.95% 1.510 ms 2.50% -1938.259 us -56.21% 🟢 FAST
I32 12 20 100000 0 50 3.051 ms 0.52% 2.546 ms 1.71% -504.582 us -16.54% 🟢 FAST
I32 12 100 100000 0 50 2.619 ms 1.04% 2.301 ms 0.66% -318.243 us -12.15% 🟢 FAST
I32 12 20 10000000 0 50 20.444 ms 0.17% 7.511 ms 0.27% -12933.268 us -63.26% 🟢 FAST
I32 12 100 10000000 0 50 19.458 ms 0.36% 5.510 ms 0.27% -13948.042 us -71.68% 🟢 FAST
I32 12 20 100000 0.5 50 3.199 ms 1.19% 2.666 ms 0.99% -532.914 us -16.66% 🟢 FAST
I32 12 100 100000 0.5 50 2.745 ms 0.73% 2.631 ms 1.76% -113.669 us -4.14% 🟢 FAST
I32 12 20 10000000 0.5 50 15.412 ms 0.50% 7.946 ms 0.81% -7466.478 us -48.44% 🟢 FAST
I32 12 100 10000000 0.5 50 14.144 ms 0.17% 5.840 ms 0.27% -8304.323 us -58.71% 🟢 FAST
F64 12 20 100000 0 1 191.239 us 2.39% 165.382 us 2.44% -25.857 us -13.52% 🟢 FAST
F64 12 100 100000 0 1 179.834 us 2.97% 162.255 us 2.89% -17.579 us -9.78% 🟢 FAST
F64 12 20 10000000 0 1 1.269 ms 0.53% 706.883 us 0.70% -562.409 us -44.31% 🟢 FAST
F64 12 100 10000000 0 1 1.137 ms 0.64% 580.814 us 0.72% -555.723 us -48.90% 🟢 FAST
F64 12 20 100000 0.5 1 197.215 us 3.55% 166.821 us 2.48% -30.394 us -15.41% 🟢 FAST
F64 12 100 100000 0.5 1 186.692 us 3.46% 166.977 us 2.70% -19.715 us -10.56% 🟢 FAST
F64 12 20 10000000 0.5 1 1.170 ms 0.61% 697.594 us 0.63% -472.659 us -40.39% 🟢 FAST
F64 12 100 10000000 0.5 1 1.030 ms 0.89% 547.405 us 1.05% -482.651 us -46.86% 🟢 FAST
F64 12 20 100000 0 10 809.880 us 2.05% 660.009 us 1.33% -149.871 us -18.51% 🟢 FAST
F64 12 100 100000 0 10 720.840 us 2.08% 649.401 us 1.11% -71.438 us -9.91% 🟢 FAST
F64 12 20 10000000 0 10 4.950 ms 0.29% 2.738 ms 0.34% -2212.160 us -44.69% 🟢 FAST
F64 12 100 10000000 0 10 4.593 ms 0.31% 2.296 ms 0.48% -2296.588 us -50.00% 🟢 FAST
F64 12 20 100000 0.5 10 839.168 us 1.55% 686.650 us 1.03% -152.518 us -18.17% 🟢 FAST
F64 12 100 100000 0.5 10 744.561 us 1.73% 699.412 us 2.10% -45.149 us -6.06% 🟢 FAST
F64 12 20 10000000 0.5 10 3.920 ms 0.25% 2.597 ms 0.39% -1322.853 us -33.74% 🟢 FAST
F64 12 100 10000000 0.5 10 3.561 ms 0.40% 1.995 ms 0.46% -1565.797 us -43.97% 🟢 FAST
F64 12 20 100000 0 50 3.499 ms 1.81% 2.845 ms 1.07% -654.225 us -18.70% 🟢 FAST
F64 12 100 100000 0 50 3.041 ms 1.93% 2.795 ms 5.15% -245.744 us -8.08% 🟢 FAST
F64 12 20 10000000 0 50 21.334 ms 0.44% 11.833 ms 0.28% -9500.360 us -44.53% 🟢 FAST
F64 12 100 10000000 0 50 19.940 ms 0.22% 10.011 ms 1.13% -9929.571 us -49.80% 🟢 FAST
F64 12 20 100000 0.5 50 3.685 ms 0.93% 3.025 ms 3.69% -660.328 us -17.92% 🟢 FAST
F64 12 100 100000 0.5 50 3.187 ms 0.90% 3.035 ms 1.82% -151.887 us -4.77% 🟢 FAST
F64 12 20 10000000 0.5 50 16.030 ms 0.15% 11.045 ms 0.36% -4984.915 us -31.10% 🟢 FAST
F64 12 100 10000000 0.5 50 14.717 ms 0.22% 8.427 ms 1.40% -6289.485 us -42.74% 🟢 FAST
groupby_m2_var_std, 32 to 1000 rows per group: 56 faster, 0 slower, 0 same
T U value_key_ratio num_rows null_probability num_aggs Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I32 12 32 10000000 0 1 1.171 ms 1.38% 579.191 us 2.65% -591.686 us -50.53% 🟢 FAST
I32 12 50 10000000 0 1 1.139 ms 0.51% 545.784 us 1.97% -593.706 us -52.10% 🟢 FAST
I32 12 64 10000000 0 1 1.133 ms 0.61% 531.030 us 1.43% -601.715 us -53.12% 🟢 FAST
I32 12 100 10000000 0 1 1.122 ms 0.62% 498.915 us 2.67% -622.868 us -55.52% 🟢 FAST
I32 12 200 10000000 0 1 1.106 ms 0.57% 473.594 us 3.66% -631.966 us -57.16% 🟢 FAST
I32 12 500 10000000 0 1 1.147 ms 0.48% 426.998 us 2.15% -720.084 us -62.78% 🟢 FAST
I32 12 1000 10000000 0 1 1.186 ms 1.33% 418.391 us 2.93% -767.481 us -64.72% 🟢 FAST
I32 12 32 10000000 0.5 1 1.073 ms 0.51% 581.830 us 0.82% -491.131 us -45.77% 🟢 FAST
I32 12 50 10000000 0.5 1 1.039 ms 1.20% 549.611 us 0.91% -489.088 us -47.09% 🟢 FAST
I32 12 64 10000000 0.5 1 1.033 ms 0.66% 537.734 us 1.52% -494.897 us -47.93% 🟢 FAST
I32 12 100 10000000 0.5 1 1.017 ms 0.72% 508.862 us 2.15% -507.880 us -49.95% 🟢 FAST
I32 12 200 10000000 0.5 1 1.005 ms 0.63% 477.605 us 1.64% -527.418 us -52.48% 🟢 FAST
I32 12 500 10000000 0.5 1 1.019 ms 0.83% 422.756 us 1.43% -595.911 us -58.50% 🟢 FAST
I32 12 1000 10000000 0.5 1 1.084 ms 0.53% 413.208 us 4.21% -671.218 us -61.90% 🟢 FAST
I32 12 32 10000000 0 10 4.541 ms 0.22% 1.618 ms 0.50% -2923.001 us -64.36% 🟢 FAST
I32 12 50 10000000 0 10 4.490 ms 0.50% 1.507 ms 0.50% -2983.150 us -66.44% 🟢 FAST
I32 12 64 10000000 0 10 4.473 ms 0.28% 1.481 ms 0.50% -2991.806 us -66.89% 🟢 FAST
I32 12 100 10000000 0 10 4.487 ms 0.19% 1.424 ms 0.66% -3062.979 us -68.26% 🟢 FAST
I32 12 200 10000000 0 10 4.597 ms 0.25% 1.527 ms 0.50% -3069.931 us -66.78% 🟢 FAST
I32 12 500 10000000 0 10 4.975 ms 0.24% 1.323 ms 0.71% -3652.037 us -73.41% 🟢 FAST
I32 12 1000 10000000 0 10 5.649 ms 0.70% 1.275 ms 1.66% -4373.617 us -77.42% 🟢 FAST
I32 12 32 10000000 0.5 10 3.567 ms 0.26% 1.686 ms 0.63% -1881.070 us -52.74% 🟢 FAST
I32 12 50 10000000 0.5 10 3.493 ms 0.34% 1.587 ms 0.60% -1905.622 us -54.56% 🟢 FAST
I32 12 64 10000000 0.5 10 3.476 ms 0.39% 1.536 ms 0.58% -1939.713 us -55.81% 🟢 FAST
I32 12 100 10000000 0.5 10 3.460 ms 0.35% 1.517 ms 2.38% -1943.391 us -56.17% 🟢 FAST
I32 12 200 10000000 0.5 10 3.505 ms 0.21% 1.567 ms 2.12% -1937.713 us -55.29% 🟢 FAST
I32 12 500 10000000 0.5 10 3.674 ms 0.79% 1.277 ms 3.53% -2396.259 us -65.23% 🟢 FAST
I32 12 1000 10000000 0.5 10 3.956 ms 0.97% 1.211 ms 1.55% -2744.483 us -69.38% 🟢 FAST
F64 12 32 10000000 0 1 1.192 ms 0.51% 647.735 us 1.94% -543.949 us -45.65% 🟢 FAST
F64 12 50 10000000 0 1 1.158 ms 1.04% 630.115 us 3.45% -527.639 us -45.57% 🟢 FAST
F64 12 64 10000000 0 1 1.149 ms 0.53% 617.447 us 1.95% -531.955 us -46.28% 🟢 FAST
F64 12 100 10000000 0 1 1.137 ms 0.39% 593.743 us 1.87% -542.822 us -47.76% 🟢 FAST
F64 12 200 10000000 0 1 1.127 ms 0.61% 588.834 us 4.15% -538.386 us -47.76% 🟢 FAST
F64 12 500 10000000 0 1 1.165 ms 0.53% 498.708 us 3.11% -666.555 us -57.20% 🟢 FAST
F64 12 1000 10000000 0 1 1.283 ms 1.41% 490.997 us 6.78% -792.079 us -61.73% 🟢 FAST
F64 12 32 10000000 0.5 1 1.087 ms 0.52% 630.771 us 2.26% -456.697 us -42.00% 🟢 FAST
F64 12 50 10000000 0.5 1 1.054 ms 0.59% 594.562 us 1.95% -459.246 us -43.58% 🟢 FAST
F64 12 64 10000000 0.5 1 1.051 ms 0.70% 579.340 us 1.86% -471.831 us -44.89% 🟢 FAST
F64 12 100 10000000 0.5 1 1.033 ms 0.79% 553.515 us 1.87% -479.795 us -46.43% 🟢 FAST
F64 12 200 10000000 0.5 1 1.023 ms 0.64% 578.885 us 1.52% -443.745 us -43.39% 🟢 FAST
F64 12 500 10000000 0.5 1 1.027 ms 0.84% 488.695 us 5.23% -538.389 us -52.42% 🟢 FAST
F64 12 1000 10000000 0.5 1 1.045 ms 0.50% 460.389 us 5.87% -584.527 us -55.94% 🟢 FAST
F64 12 32 10000000 0 10 4.678 ms 0.25% 2.428 ms 2.30% -2249.312 us -48.09% 🟢 FAST
F64 12 50 10000000 0 10 4.627 ms 0.18% 2.354 ms 1.89% -2272.400 us -49.12% 🟢 FAST
F64 12 64 10000000 0 10 4.603 ms 0.22% 2.328 ms 0.53% -2275.034 us -49.43% 🟢 FAST
F64 12 100 10000000 0 10 4.623 ms 0.21% 2.302 ms 0.33% -2321.549 us -50.21% 🟢 FAST
F64 12 200 10000000 0 10 4.680 ms 0.17% 2.667 ms 0.28% -2013.148 us -43.02% 🟢 FAST
F64 12 500 10000000 0 10 5.155 ms 0.19% 2.013 ms 0.46% -3142.420 us -60.96% 🟢 FAST
F64 12 1000 10000000 0 10 5.779 ms 0.46% 1.970 ms 0.37% -3809.054 us -65.91% 🟢 FAST
F64 12 32 10000000 0.5 10 3.671 ms 0.24% 2.181 ms 0.40% -1489.885 us -40.59% 🟢 FAST
F64 12 50 10000000 0.5 10 3.614 ms 0.26% 2.045 ms 0.36% -1568.910 us -43.41% 🟢 FAST
F64 12 64 10000000 0.5 10 3.600 ms 0.31% 2.033 ms 0.43% -1567.141 us -43.53% 🟢 FAST
F64 12 100 10000000 0.5 10 3.584 ms 0.31% 2.002 ms 0.39% -1582.186 us -44.14% 🟢 FAST
F64 12 200 10000000 0.5 10 3.599 ms 0.76% 2.642 ms 0.32% -956.691 us -26.58% 🟢 FAST
F64 12 500 10000000 0.5 10 3.835 ms 2.04% 1.836 ms 0.39% -1999.688 us -52.14% 🟢 FAST
F64 12 1000 10000000 0.5 10 4.082 ms 0.24% 1.614 ms 0.48% -2467.651 us -60.45% 🟢 FAST
sum: 7 faster, 1 slower, 0 same
T num_rows Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
I64 100000 138.087 us 4.84% 118.344 us 7.06% -19.743 us -14.30% 🟢 FAST
I64 1000000 201.953 us 2.90% 144.028 us 10.34% -57.925 us -28.68% 🟢 FAST
I64 10000000 794.452 us 4.45% 532.425 us 4.01% -262.027 us -32.98% 🟢 FAST
I64 100000000 5.519 ms 0.50% 4.441 ms 0.16% -1077.376 us -19.52% 🟢 FAST
decimal64 100000 142.626 us 6.32% 126.947 us 3.13% -15.679 us -10.99% 🟢 FAST
decimal64 1000000 346.463 us 1.59% 240.369 us 1.79% -106.094 us -30.62% 🟢 FAST
decimal64 10000000 2.378 ms 0.41% 1.763 ms 0.24% -614.597 us -25.84% 🟢 FAST
decimal64 100000000 26.281 ms 0.11% 29.731 ms 0.13% 3.451 ms 13.13% 🔴 SLOW
no_requests: 3 faster, 0 slower, 1 same
num_rows Ref Time Ref Noise Cmp Time Cmp Noise Diff %Diff Status
100000 48.707 us 14.64% 48.157 us 10.98% -0.550 us -1.13% 🔵 SAME
1000000 94.691 us 9.48% 56.921 us 3.33% -37.770 us -39.89% 🟢 FAST
10000000 336.131 us 1.01% 173.423 us 2.54% -162.708 us -48.41% 🟢 FAST
100000000 2.826 ms 0.29% 917.864 us 0.61% -1907.940 us -67.52% 🟢 FAST

Use device-scope table atomics and preserve overflow, nullable, and
all-unique grouping invariants. Bound group counters by the input size
when the table is larger, and release the table before compaction.

Split reduction instantiations across translation units with shared CUB
dispatch. Preserve streaming aggregation support and device-view ownership.
Add regression coverage for the corrected grouping and streaming behavior.

PointKernel commented Sep 10, 2026

Copy link
Copy Markdown
Member Author

Follow-up to the previous benchmark report, at commit ea84e5c09d. RTX PRO 6000 (SM120), CUDA 12.9 / CCCL 3.5, idle GPU. The final benchmark-skill sweep reran the candidate and reused saved baselines: PR 11d3dd1998 and main 2062f4d3241.

Major changes since 11d3dd1998:

  • Fixed atomic slot access, retry/overflow handling, large-input bounds, nullable reductions, streaming support checks, and cached device-view ownership.
  • Bound group counters by min(rows, table capacity), reuse representative row IDs already returned by insertion, and release the hash table before compaction. The all-unique path skips unnecessary grouped-row construction.
  • Split reducer compilation across six translation units, sharing CUB dispatch and giving each explicit instantiation one owner. Reused existing CCCL/CUB, Thrust, cuDF and RMM facilities; no new public APIs or GPU-specific tuning/dispatch.

GPU runtime change is the geometric mean of per-case time ratios; negative means faster. The 330 official cases include 16 additional small-input cases beyond the previous report.

Benchmark Cases vs saved main vs previous PR
groupby_max_cardinality 100 -83.60% +1.08%
complex_int_keys 20 -57.78% +2.21%
complex_mixed_keys 40 -28.65% -12.84%
groupby_max 36 -50.87% -41.51%
groupby_struct_keys 18 -5.01% +6.40%
groupby_m2_var_std 104 -45.04% +2.90%
sum 8 -33.76% -15.59%
no_requests 4 -40.04% +6.97%
Total 330 -60.39% -5.97%

The median case is 1.62% slower than the previous PR; the largest slowdown is 0.286 ms. Every measured GPU and synchronized CPU comparison against PR/main satisfies the agreed ≤5% slowdown or <1 ms increase.

Selected cases use milliseconds and decimal MB. Peak is requested device memory, listed as main / previous PR / now.

Case Main ms Previous PR ms Now ms Peak MB: main / PR / now
20M rows, 8 groups, 8 I32 MAX 1.720 1.898 1.901 400.6 / 321.0 / 241.0
20M rows, 1M groups, 8 I32 MAX 4.791 3.050 2.948 560.6 / 339.0 / 259.0
groupby_max: 16.8M rows, 32 I32 MAX, 90% value nulls 7.202 7.657 5.378 718.5 / 519.0 / 507.4
100M rows, decimal64 SUM 26.572 29.746 15.157 2552.3 / 2875.8 / 1773.2
16.8M rows, 3 mixed keys, value_key_ratio=200, 50% nulls 2.606 1.859 1.556 416.4 / 476.3 / 337.6
16.8M all-unique stress 7.991 11.199 4.165 469.8 / 536.9 / 469.8

Across 333 official + all-unique cases, peak memory is lower than saved main in 328, equal in 3, and higher by just 500 bytes in each of two keys-only cases. None exceed the previous PR. Eight streaming cases also passed against the saved local reference with identical peaks; no saved upstream streaming baseline is available.

Uncached SM120 compilation: grouping 35.3 s, longest measured TU 132.6 s across 17 production and two test objects, reusing earlier timings for unchanged objects. The 150 s/TU limit is not established for multi-architecture builds. libcudf.so is 233.31 MB, down 7.02 MB from the PR but still 8.28 MB (+3.68%) above saved main.

Whole-PR self-review and applicable pre-commit hooks passed, along with 119 C++ suites, 25 Python groupby modules, and 21 Compute Sanitizer memcheck tests with zero errors.

Preserve concurrent streaming aggregation and the cached device-view
ownership fix while resolving the overlapping standard-library includes.
if (!is_nested(input.type())) { return 1; }
auto const requested =
std::max(static_cast<double>(num_rows) + 1,
std::ceil(static_cast<double>(num_rows) / cudf::detail::CUCO_DESIRED_LOAD_FACTOR));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is 0.5 load, but I believe we discussed that higher load factors were still viable with good performance for HashCSR.

Equal const& d_row_equal,
Hash const& d_row_hash,
bool need_grouped_rows,
cuda::stream_ref stream)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should take a cudf::memory_resources for temp MR control.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2 - In Progress Currently a work in progress CMake CMake build issue improvement Improvement / enhancement to an existing function libcudf Affects libcudf (C++/CUDA) code. non-breaking Non-breaking change Performance Performance related issue

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants